Papers with Closed-loop mitigation
Detecting Proxy Gaming in RL and LLM Alignment via Evaluator Stress Tests (2026.findings-acl)
Copied to clipboard
| Challenge: | Proxy optimization is a challenge spanning reinforcement learning and LLM alignment. |
| Approach: | They propose an invariance-based framework that detects proxy gaming by separating exploitable sensitivity from content-driven improvements using semantic validity audits. |
| Outcome: | The proposed framework achieves 78.4% precision and 81.7% recall across 15 environments and 5 algorithms. |